Papers with corpus annotation

6 papers
A Simple yet Efficient Prompt Compression Method for Text Classification Data Annotation Using LLM (2025.coling-industry)

Copied to clipboard

Challenge: Existing methods to improve the accuracy of large language models (LLMs) are often impractical due to high costs and time consumption.
Approach: They propose a method that uses keyword extraction to reduce prompt tokens in text annotation tasks.
Outcome: The proposed method reduces prompt tokens while maintaining high accuracy.
The ISO Standard for Dialogue Act Annotation, Second Edition (2020.lrec-1)

Copied to clipboard

Challenge: ISO standard 24617-2 for dialogue act annotation has been used in corpus annotation and in the design of components for spoken and multimodal interactive systems.
Approach: ISO standard 24617-2 for dialogue act annotation is proposed for a second edition . this second edition allows a more accurate annotation of dependence relations and rhetorical relations in dialogue.
Outcome: The proposed second edition of ISO 24617-2 for dialogue act annotation addresses some inaccuracies and undesirable limitations.
Moving TIGER beyond Sentence-Level (L18-1)

Copied to clipboard

Challenge: TIGER 2.2-doc is a new set of annotations for the German TIger corpus.
Approach: They propose a new set of annotations for the German TIGER corpus . they introduce new document-level annotations: authors and their gender.
Outcome: The new annotations improve the TIGER corpus and its structure and authors and gender.
Arabic Speech Rhythm Corpus: Read and Spontaneous Speaking Styles (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of Arabic speech recordings has been built to allow comparisons between Arabic and other languages.
Approach: They propose to build a corpus of Arabic speech recordings that can be compared with other languages.
Outcome: The proposed corpus can be used for forensic phonetic research and casework applications.
Parallel Corpora in Mboshi (Bantu C25, Congo-Brazzaville) (L18-1)

Copied to clipboard

Challenge: BULB project aims to provide tools to language documentation and description for unwritten languages . language-based technologies are needed to support the collection of data and to provide linguistic documentation for the languages.
Approach: This paper presents multimodal and parallel data collections in Mboshi, as part of the French-German BULB project.
Outcome: The proposed data collection includes pictures and videos documenting social practices, agriculture, wildlife and plants.
Polish Discourse Corpus (PDC): Corpus Design, ISO-Compliant Annotation, Data Highlights, and Parser Development (2024.lrec-main)

Copied to clipboard

Challenge: The Polish Discourse Corpus employs ISO 24617-8 for discourse relation annotation.
Approach: They propose to adopt ISO 24617-8 standard for discourse relation annotation for Polish and to develop a parser tailored for the framework.
Outcome: The Polish Discourse Corpus adopts ISO 24617-8, a segment of the Language Resource Management – Semantic Annotation Framework (SemAF) the paper examines the corpus architecture, annotation procedures, and the challenges encountered by annotators.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations